Skip to content

feat(launch): re-add llm-d multinode support to infx.launch - #3613

Merged
adibarra merged 8 commits into
mainfrom
feat/llmd-python-launcher
Oct 2, 2026
Merged

adibarra merged 8 commits into
mainfrom
feat/llmd-python-launcher

Conversation

@adibarra

@adibarra adibarra commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Re-adds llm-d (llmd-vllm) multinode support on the Python launcher from #3576, which dropped the launch_gb200-nv.sh llm-d branch.

llm-d keeps its own Slurm orchestration (submit.sh → sbatch job.slurm → srun + pyxis per node → server.sh). The launcher owns everything around it:

  • policy.launch_path: multinode framework == llmd-vllm → LaunchPath.LLMD, enabled on gb200-nv only (LLMD_CLUSTERS); the table check requires slurm.squash.
  • drivers/llmd.py: resolves the checkpoint from runners.yaml, imports the image via backend.prepare_image, runs llm-d/submit.sh directly with the point's topology, sequence lengths and concurrency, attaches to the printed job id, streams the log, checks the final Slurm state, and stages result JSONs, agentic and eval artifacts plus the server-log bundle. No per-model wrapper script.
  • srt/models.py: dsv4 / fp4 / llmd-vllm on gb200-nv → DeepSeek-V4-Pro@numa1, served as deepseek-ai/DeepSeek-V4-Pro.

Fix: successful jobs no longer end CANCELLED

job.slurm used to scancel its own allocation once the decode coordinator finished, so every successful run ended CANCELLED 0:0, which the launcher reports as a failure (run 36720223235 finished its benchmark and was still marked failed). job.slurm now stops the srun step and exits 0, so the job ends COMPLETED. A step that dies before the coordinator finishes still fails the job with its exit code.

Driver port by @ilmarkov from #2719.

Testing

  • pytest infx/tests/launch infx/tests/clusters infx/tests/matrix infx/tests/workflows: 1230 passed; ruff clean
  • validate_perf_changelog against origin/main: passes
  • Local harness for the job.slurm step control: hung workers + done marker → exit 0; step failing first → its exit code
  • GB200 e2e (with a temporary dsv4-fp4-gb200-llmd-vllm 8k1k c1 point, since removed): https://github.com/SemiAnalysisAI/InferenceX/actions/runs/36906427804 passed; the job ended COMPLETED and the result JSON was staged

Add an llm-d driver that submits the job through the llm-d wrapper, attaches
to the Slurm job, streams its log, checks its final state and stages results,
agentic and eval artifacts. Route llmd-vllm multinode requests to it on
gb200-nv and resolve DeepSeek-V4-Pro to the node-local numa1 checkpoint.

Restore the GB200 DeepSeek-V4-Pro disagg wrapper.

job.slurm no longer scancels its own allocation when the coordinator finishes;
it stops the srun step and exits 0, so the job ends COMPLETED instead of
CANCELLED, which infx.launch reports as a failure.

Driver port by Ilya Markov from #2719.
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

…est point

Drop the per-model llm-d wrapper script. The driver passes topology, sequence
lengths and concurrency to submit.sh and sets GPUS_PER_NODE, TIME_LIMIT,
CONTAINER_IMAGE and worker counts itself. llmd-vllm supports only P/D
disaggregated points.

Re-add the dsv4-fp4-gb200-llmd-vllm low-latency 8k1k point (1P DEP8 + 1D TP8,
conc 1) to exercise the path end to end.
Route every multinode llmd-vllm request to the llm-d driver. The driver
requires slurm.squash and a staged checkpoint, which is what the whitelist
stood in for.
/mnt/numa1 exists only on compute nodes, so the runner host cannot see the
checkpoint. job.slurm already checks it on every allocated node.
…cher

# Conflicts:
#	inferencex-e2e/infx/launch/drivers/__init__.py
#	inferencex-e2e/infx/launch/policy.py
#	inferencex-e2e/infx/launch/request.py
#	inferencex-e2e/perf-changelog.yaml
…cher

# Conflicts:
#	inferencex-e2e/perf-changelog.yaml
The point validated the llm-d path end to end in e2e run 36906427804; it is not
a benchmark to publish.
@adibarra
adibarra marked this pull request as ready for review October 2, 2026 21:57
@adibarra
adibarra requested a review from a team October 2, 2026 21:57
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @adibarra's task in 56s —— View job


Review: PR #3613

LGTM - no blocking issues found

One note, not blocking: the PR description says the path is limited to gb200-nv through LLMD_CLUSTERS, and that a table check requires slurm.squash. Neither exists in the diff. policy.launch_path (infx/launch/policy.py:80) sends llmd-vllm multinode points to LaunchPath.LLMD on any cluster, and slurm.squash is only checked when the driver runs (infx/launch/drivers/llmd.py:49). No master config uses llmd-vllm right now, so nothing is affected today. Please either add the gate or update the description.

@adibarra
adibarra merged commit d2cd310 into main Oct 2, 2026
9 of 10 checks passed
@adibarra
adibarra deleted the feat/llmd-python-launcher branch October 2, 2026 22:03
cquil11 added a commit that referenced this pull request Oct 2, 2026
Add aggregate and disaggregated AgentX recipes on the Python llm-d
launcher from #3613. Preserve explicit checkpoint selection, worker
topology, offload metadata, and per-rank metrics collection.

Keep successful Slurm completion distinct from failed or premature exits.

Co-authored-by: Cam Quilici <[email protected]>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline finding, I checked two other things the PR description's claims hinge on: policy.launch_path (inferencex-e2e/infx/launch/policy.py:80) actually routes any multinode llmd-vllm request to LaunchPath.LLMD regardless of cluster, not just gb200-nv — but srt/models.py only defines an llmd-vllm checkpoint entry for gb200-nv, so other clusters fail safely with a clear "stages no checkpoint" error rather than misbehaving. I also checked that drivers/llmd.py stages the image via backend.prepare_image() without calling run_container() — that's intentional, since submit.sh/job.slurm consume the squash reference directly through srun --container-image/pyxis rather than the launcher's normal container-run path.

Extended reasoning...

This PR re-adds llm-d multinode Slurm orchestration to the Python launcher (policy routing, a new driver, model/checkpoint resolution, and the job.slurm/server.sh scripts it shells out to) with no auth/crypto surface but real job-success/failure semantics at stake. A confirmed inline finding shows the job.slurm fix for the CANCELLED-on-success bug can itself mask a real failure as success when the decode coordinator exits early, because server.sh's EXIT trap touches the done marker unconditionally. I additionally verified two PR-description claims that looked suspicious in isolation (cluster gating in policy.py, and image staging without run_container in llmd.py) and found both fail safe rather than being functional bugs.

Comment on lines +70 to +84
run_until_bench_done() {
"$@" &
local step=$! rc=0
while kill -0 "$step" 2>/dev/null; do
if [[ -f "$BENCH_DONE_MARKER" ]]; then
echo "[${SLURM_JOB_ID}] coordinator finished; stopping the srun step"
kill -TERM "$step" 2>/dev/null || true
break
fi
sleep 5
done
wait "$step" || rc=$?
[[ -f "$BENCH_DONE_MARKER" ]] && return 0
echo "[${SLURM_JOB_ID}] srun step exited rc=$rc before the coordinator finished" >&2
return "$rc"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Any startup failure in the decode coordinator (EPP/Envoy not ready, prefill health timeout) now makes the whole llm-d job report success instead of failure, with no benchmark results. server.sh:303 has trap 'touch "$BENCH_DONE_MARKER"' EXIT, which fires on every exit of that branch, including the exit 1 failure paths, writing the same marker file as the real success touch at server.sh:573. job.slurm:82 [[ -f "$BENCH_DONE_MARKER" ]] && return 0 treats marker-exists as unconditional success, ignoring rc. Fix: use a distinct marker/exit path for genuine completion so job.slurm only succeeds when the benchmark actually ran, not merely when the coordinator exited for any reason.

Why this was flagged

Trigger: the decode-leader branch in server.sh (ROLE==decode && LWS_WORKER_INDEX==0) hits any of its exit 1 checks - EPP not binding within 60s (server.sh:384/388), Envoy /ready failing within 120s (server.sh:407/412), or prefill /health timing out within 300s (server.sh:509) - reached via job.slurm's run_until_bench_done -> srun -> server.sh. The EXIT trap at server.sh:303 touches $BENCH_DONE_MARKER on that exit, same file as the genuine-success touch at server.sh:573. job.slurm's run_until_bench_done (job.slurm:70-85) only checks file existence at line 82, not the step's rc, so it returns 0. set -eo pipefail (job.slurm:10) then lets the script finish normally, Slurm records COMPLETED, and drivers/llmd.py:137-138 computes rc=0 from that state; the result-JSON glob (llmd.py:140) finds nothing but the loop simply doesn't run, so no error is raised. On the base branch this same failure ends CANCELLED, which the launcher already treats as failure.

Verification: server.sh:303's trap 'touch "$BENCH_DONE_MARKER"' EXIT fires on every exit of the decode-coordinator branch, including the startup exit 1 failures at server.sh:384, 388, 407, 412, and 509, writing the same marker as the genuine-success touch at server.sh:573. job.slurm's [[ -f "$BENCH_DONE_MARKER" ]] && return 0 then discards the failing rc whenever that marker exists.

cquil11 added a commit that referenced this pull request Oct 2, 2026
Add aggregate and disaggregated AgentX recipes on the Python llm-d
launcher from #3613. Preserve explicit checkpoint selection, worker
topology, offload metadata, and per-rank metrics collection.

Keep successful Slurm completion distinct from failed or premature exits.
Use golden synthetic acceptance for every AgentX MTP throughput role and
launch accuracy evaluations separately with real verification.

Co-authored-by: Cam Quilici <[email protected]>
chunfangamd added a commit that referenced this pull request Oct 2, 2026
#3613 routes multi-node llmd-vllm points to the llm-d driver instead of
srt-slurm, so they name no srt-slurm recipe and need no CONFIG_FILE. The
preflight now checks only the points the launcher hands to srt-slurm,
and both read the same LLMD_FRAMEWORKS set from infx.launch.policy.

#3613 让多节点的 llmd-vllm 点改由 llm-d driver 启动,而不是 srt-slurm,
因此这些点不引用 srt-slurm 配方,也不需要 CONFIG_FILE。预检现在只检查
launcher 交给 srt-slurm 的点,二者共用 infx.launch.policy 中的
LLMD_FRAMEWORKS。

Co-authored-by: Cursor <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant